Japanese TN fix and improvement - #444
Conversation
Signed-off-by: Mai Anh <palasek182@gmail.com>
for more information, see https://pre-commit.ci
|
please update jenkinsfile |
| Finite state transducer for classifying Japanese address-like expressions. | ||
|
|
||
| Examples: | ||
| 東京都千代田区丸の内1-1-1 -> name: "東京都千代田区丸の内一の一の一" |
There was a problem hiding this comment.
let's make sure all tagger outputs are fully formatted properly
| graph_integer = ( | ||
| pynutil.insert('integer_part: \"') | ||
| + ( | ||
| pynini.cross("0", "零") |
There was a problem hiding this comment.
wouldn't this be zero_decimal?
| def __init__(self, deterministic: bool = True): | ||
| super().__init__(name="punctuation", kind="classify", deterministic=deterministic) | ||
| s = "!#$%&'()*+,-./:;<=>?@^_`{|}。,;:《》“”·~【】!?、‘’.<>-——_、。.「」『』‘`/・;’”“”‷・〔〕々〃ゝゞヽ〲〱〳〴〵ヾ〆,~" | ||
| s = "!#$%&'()*+,-./:;<=>?@^_`{|}~" |
There was a problem hiding this comment.
are we sure we don't need this punctuation?
| division = pynini.string_file(get_abs_path("data/time/division.tsv")) | ||
|
|
||
| division_component = pynutil.insert("suffix: \"") + division + pynutil.insert("\"") | ||
| hour_number = pynutil.add_weight(pynini.cross("0", "零"), -0.1) | graph_cardinal |
There was a problem hiding this comment.
again looks like zero_decimal?
Signed-off-by: Mai Anh <palasek182@gmail.com>
for more information, see https://pre-commit.ci
Signed-off-by: Mai Anh <palasek182@gmail.com>
|
This PR is stale because it has been open for 14 days with no activity. Remove stale label or comment or update or this will be closed in 7 days. |
|
This PR was closed because it has been inactive for 7 days since being marked as stale. |
|
This PR is stale because it has been open for 14 days with no activity. Remove stale label or comment or update or this will be closed in 7 days. |
| class TestTelephoneExtended: | ||
| normalizer_ja = Normalizer(lang='ja', cache_dir=CACHE_DIR, overwrite_cache=False, input_case='cased') | ||
|
|
||
| @parameterized.expand(parse_test_case_file('ja/data_text_normalization/test_cases_telephone_extended.txt')) |
There was a problem hiding this comment.
why is this a separate file? and would it also need to be added to sh tests?
| @@ -0,0 +1,14 @@ | |||
| B2A23C~ビー 二 エー 二三 シー | |||
There was a problem hiding this comment.
let's double check if this is the preferred method for the language in TTS -- it may still be useful to have the English letters while normalizing numbers
| 0.05m | ||
| -> measure { decimal { integer_part: "零" fractional_part: "零五" } units: "メートル" preserve_order: true } | ||
| 60km/h -> measure { cardinal { integer: "時速六十" } units: "キロ" preserve_order: true } | ||
| 50m/s -> measure { cardinal { integer: "秒速五十" } units: "メートル" preserve_order: true } |
There was a problem hiding this comment.
i'm seeing cardinal integer text containing the unit here and the unit being incorrectly defined? is it just the example or an implementation error?
| Finite state transducer for classifying Japanese ranges. | ||
|
|
||
| Examples: | ||
| 2-5 -> name: "二から五" |
| Finite state transducer for classifying Roman numerals in supported contexts. | ||
|
|
||
| Examples: | ||
| 第III章 -> name: "第三章" |
What does this PR do ?
Add a one line overview of what this PR aims to accomplish.
Before your PR is "Ready for review"
Pre checks:
git commit -sto sign.pytestor (if your machine does not have GPU)pytest --cpufrom the root folder (given you marked your test cases accordingly@pytest.mark.run_only_on('CPU')).bash tools/text_processing_deployment/export_grammars.sh --MODE=test ...pytestand Sparrowhawk here.__init__.pyfor every folder and subfolder, includingdatafolder which has .TSV files?Copyright (c) 2023, NVIDIA CORPORATION & AFFILIATES. All rights reserved.to all newly added Python files?Copyright 2015 and onwards Google, Inc.. See an example here.try import: ... except: ...) if not already done.PR Type:
If you haven't finished some of the above items you can still open "Draft" PR.